rdmbair15m5 - changelog - agent-fleet-self-healing-and-4h-stability-review - claude - fleet-automation - agent-coordination - 20260823-1825
Session: 2026-08-22 21:31 EDT -> 2026-08-23 18:26
EDT (~21 h) Ran on: rdmbair15m5 ·
Changed state on: all six fleet hosts, canonical work
on rdmsm4x
Built an automated agent-fleet management layer: heartbeats, automatic bus sync, self-healing, independent off-host auditing with remote repair, a status board, and a 4-hourly stability review graded on delivery ratio. Rich's stated goal was to stop managing agents by hand — "it is enough just directing work, vs. managing and checking all agent work."
What now runs unattended
| Job | Cadence | Scope | Purpose |
|---|---|---|---|
com.eastcoastscience.agentstatus |
5 min | all 6 | agent heartbeat + automatic bus sync |
com.eastcoastscience.agentheal |
10 min | all 5 | self-heal wedged launchd; escalate what it cannot |
com.eastcoastscience.offhostwatch |
10 min | rdmbair15m5 | watch rdmsm4x from outside |
com.eastcoastscience.fleetaudit |
15 min | rdmsm4x | independent verification + remote healing |
com.eastcoastscience.stabilityreview |
4 h | rdmsm4x | delivery-ratio grading |
All carry RunAtLoad=true. Scripts:
~/.agent-coordination/{agent_status,agent_heal}.zsh,
~/dev/_ops/fleetwatch/offhost_watch.zsh,
rdmsm4x:~/dev/scripts/{fleet_audit,fleet_stability_review}.zsh.
The two gaps that existed before this
- The bus never synced on its own. An inventory of
every LaunchAgent on all five hosts found no job
anywhere running
agent_msg.zsh sync— including on rdmsm4x, the hub. It ran only when an agent remembered to type it. - Nothing detected a running-but-silent agent. No
liveness signal existed at all.
rdmpw3275msent its first-ever bus message at 03:34:53 once this shipped.
Agent status board
rdmsm4x:~/dev/scripts/gen_agent_board.py ->
~/dataroo.net/wiki/agents.html, wired into
publish_site.zsh and the gen_site.py nav. Live
at https://dev.dataroo.net/agents.html.
dev.dataroo.net — two fixes
- 503 root cause:
set_real_ip_from 172.16.0.0/12did not contain cloudflared's actual address192.168.107.3, soCF-Connecting-IPwas never honoured andlimit_req_zone $binary_remote_addrkeyed on the tunnel — the entire internet shared one 15 r/s bucket. Fixed by claude@rdmsm4x with the correct CIDR pluslimit_req_status 429, so throttling stops masquerading as an outage. - Auto-auth:
satisfy any+ allow-list for Rich's home IPv4/IPv6, password kept as fallback. This is only safe because real_ip is now correct — if that regresses the allow-list silently becomes allow-all. Highest-value latent hazard on the fleet.
Config normalization
Fixed 5 contradictions; verified 5 others as already-correct and
deliberately did not churn them. Record:
rdmsm4x:~/dev/fleet/NORMALIZATION-20260823.md. Notably
FLEET.md pointed at a project path that exists on no host
and contradicted CLAUDE.md; the
fleet-notes-publish skill named two scripts that do not
exist and described the retired per-host Notes folders. Added
rule 28 to AGENT_COORDINATION.md (stay on
task, route unrelated work to claude@rdmsm4x), propagated fleet-wide.
Created 11 templates at rdmsm4x:~/dev/lib/templates/.
First stability review — DEGRADED, fleet delivery 75%
rdmpw3265m 100% rdmpw3275m 100% rdmbair13m5 95% jdmbair13m5 93%
rdmsm4x 60% rdmbair15m5 2% <- FAILING, 4 wedge events
Second review, 17:22, four hours later, unattended — DEGRADED, fleet 72%
rdmpw3265m 99% rdmpw3275m 99% jdmbair13m5 93% rdmbair13m5 unreachable
rdmsm4x 59% rdmbair15m5 11% <- still FAILING
rdmbair15m5 improved 2% -> 11% over five hours WITHOUT a reboot:
19 of ~164 expected runs. The remote heal from rdmsm4x keeps rescuing
it; launchd keeps refusing to run it unaided. That is the distinction
the metric exists to make — the host is being kept alive FROM OUTSIDE,
not recovering. Its three jobs read ok with runs 23/22/25,
which is exactly the misleading snapshot the delivery ratio sees
through. The 4 h cycle itself is proven: it ran on time,
unattended, twice.
The finding that changed the design
Self-healing cannot fix a broken launchd.
agentheal was itself wedged by the exact condition it
repairs — launchd is what spawns the healer. Verified: at 12:50 it
healed two jobs; by 13:17 all three including itself were
pended. Only remote healing from rdmsm4x recovered
it. Cadence raised 3600s -> 900s accordingly; a full audit
costs 7 seconds.
Method note
Eighteen defects across five agents this session, every one
the same shape: a proxy standing in for the thing. mtime for
write-time · transcript text for execution · exit status for data health
· tail's status for the command's · pgrep -f
for a running program · a monitor's log for the monitor running ·
state = not running for a job that is stuck pending · host
uptime for how long a job should have been running. Measure the
thing itself. Four-plus were caught only because another agent
checked the first one's work.
Outstanding for Rich
- Reboot — in progress. rdmbair15m5's launchd is wedged and self-repair provably cannot hold it.
- LogTTY authoritative tree · DevSort commit authority · the unidentified LogTTY copier — all parked.
devsite60% /fleethealth77% — likely a benign cadence overrun, left with claude@rdmsm4x.- Distribution certificates: zero on the fleet. Notarization and App Store blocked.
Verification
All job states read via launchctl print
(runs, last exit code,
pended nondemand spawn) rather than from output files — a
heartbeat file is exactly what a hand-run also produces.
Secrets
None written. Scripts read credentials at runtime from
~/.secrets/global.env.